Optimization method for data collection information age in wireless rechargeable sensor networks

Through the weighted K-Medoids clustering and DDQN reinforcement learning algorithm, the drone path and hover points are optimized, and the energy limitation and path planning conflicts in the drone wireless charging network are solved, efficient and flexible data acquisition is achieved, adapting to dynamic environmental changes, and data real-time and network performance are improved.

CN120358506BActive Publication Date: 2025-08-22NANJING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510849373.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-08-22
Estimated Expiration
2045-06-24

AI Technical Summary

Technical Problem

There are problems in the drone wireless charging network with energy limitation, conflicts between path planning and information age optimization, and poor adaptability of dynamic environments, which makes it difficult to ensure the sustainability and reliability of data acquisition.

Method used

Weighted K-Medoids clustering and dual deep Q network (DDQN) reinforcement learning algorithm are adopted to dynamically optimize the drone path and hover point layout, and combined with the weight adjustment of sensor nodes, real-time and adaptive optimization of data acquisition are achieved.

Benefits of technology

It improves the execution efficiency and flexibility of data acquisition tasks, reduces data latency, enhances the adaptability and scalability of the system, and ensures real-time information and network performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120358506B_ABST
    Figure CN120358506B_ABST
Patent Text Reader

Abstract

This application, which belongs to the field of Internet of Things technology and reinforcement learning, discloses a method for optimizing the age of information collected in wireless rechargeable sensor networks. The method first uses a weighted K-Medoids clustering algorithm to determine the drone's hovering point and the initial clustering of nodes based on the sensor nodes' geographic location, power status, and data priority. The method then uses the Dual Deep Q Network (DDQN) algorithm in deep reinforcement learning to plan the drone's path and dynamically adjust its hovering strategy. The optimization objectives are to minimize the average age of information (AoI) within a cluster and the average AoI of the entire network. The DDQN algorithm's feedback mechanism dynamically adjusts node weights in the clustering algorithm to achieve continuous optimization of the AoI. This application effectively reduces data collection latency, improves the real-time nature of information collection, and improves overall network performance. It offers the advantages of high efficiency, flexibility, adaptability, and scalability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of Internet of Things technology and reinforcement learning, and specifically relates to a method for optimizing the age of data collection information in a wireless rechargeable sensor network. Background Art

[0002] In recent years, unmanned aerial vehicle (UAV) and wireless rechargeable sensor network (WRSN) technologies have rapidly developed and are widely used in the Internet of Things (IoT) and smart city sectors. WRSNs use wireless energy transmission to provide continuous energy to sensor nodes, while drones, with their high maneuverability, flexibility, and wide coverage, have become a crucial tool for IoT data collection. The integration of these two technologies offers promising solutions for scenarios such as remote monitoring, environmental perception, and the agricultural IoT.

[0003] In IoT applications, real-time data transmission is crucial. To measure data freshness and real-time performance, the Age of Information (AoI) metric has been proposed. This metric measures the time interval between data generation and successful collection. A smaller AoI value indicates fresher data and a greater ability to meet real-time requirements.

[0004] However, the actual deployment of UAV wireless charging networks still faces the following problems: (1) Energy limitation problem: Sensor nodes (SN) usually rely on wireless energy transmission technology for energy replenishment, and the endurance of the UAV itself is also limited by the battery capacity. This dual energy constraint limits the continuity and reliability of data collection, and how to effectively manage energy becomes a major challenge. (2) Path planning and AoI optimization problem: UAVs need to consider multiple goals when performing tasks, such as maximizing the coverage area and minimizing flight energy consumption, but multi-objective planning usually conflicts with the optimization goal of AoI, making it difficult to take both into account, resulting in difficulty in ensuring the real-time transmission of data. (3) Dynamic environment adaptability problem: In the actual WRSN environment, the location, number, status, such as power level, of sensor nodes may change randomly, and some nodes may even malfunction or fail, resulting in the inability of traditional static path planning algorithms to effectively cope with it, seriously affecting the reliability of data collection. Therefore, there is an urgent need for a UAV path planning method that can dynamically adapt to environmental changes and take into account both energy constraints and real-time performance, namely AoI optimization. Summary of the Invention

[0005] In order to solve the above technical problems, this application provides a method for optimizing the age of data collection information in a wireless rechargeable sensor network. The method is based on weighted K-Medoids clustering and reinforcement learning DDQN algorithm, aiming to optimize the real-time performance of data collection in wireless rechargeable sensor networks (WRSNs), reduce the overall system age of information (AoI), and improve the data collection performance of UAV wireless charging networks.

[0006] In order to achieve the above objectives, this application is implemented through the following technical solutions:

[0007] This application is a method for optimizing the age of data collection information in a wireless rechargeable sensor network, which specifically includes the following steps:

[0008] Step 1: Build a network architecture: The network architecture is a wireless charging network for drones (WRSN), which includes multiple sensor nodes (SN) and a data center (DC). Drones depart from the DC to perform data collection and wireless charging tasks, and return to the DC after the tasks are completed.

[0009] Step 2: Initial hovering point (HP) selection: The weighted K-Medoids clustering algorithm is used to divide the sensor nodes in the wireless charging network (WRSN) of unmanned aerial vehicles into multiple clusters. The center point of each cluster is used as the initial hovering point (HP) of the drone. The initial hovering point layout and node division results of the drone are finally determined.

[0010] Step 3: Using the initial hovering point (HP) layout and node division results of the drone determined in step 2 as the initial layout, use the double depth The network (DDQN) reinforcement learning algorithm dynamically optimizes the drone's data collection path, initial hovering point (HP) visit sequence, and hovering time, enabling the drone to collect data and charge each cluster node;

[0011] Step 4: After each charging task is executed, according to the double depth The feedback results of the network (DDQN) reinforcement learning algorithm dynamically adjust the weights of the nodes in the weighted K-Medoids clustering algorithm, so that the weights of the sensor nodes in the weighted K-Medoids clustering algorithm are proportional to the current age of information (AoI) of the sensor nodes, achieving dynamic optimization of the average age of information (AoI) within the cluster and the overall average age of information (AoI) of the wireless unmanned aerial vehicle charging network (WRSN);

[0012] Step 5: Repeat steps 3 and 4 to continue the drone data collection and charging tasks until the average age of information (AoI) of the drone wireless charging network (WRSN) reaches a stable optimization state or reaches the preset maximum number of optimization iterations.

[0013] A further improvement of the present invention is that: the step 2 uses the weighted K-Medoids clustering algorithm to determine the initial hovering point layout and node division results of the drone, which specifically includes the following steps:

[0014] Step 2.1, Initialize network parameters: Set the coverage radius of the drone to , there are a total of sensor nodes , each of the sensor nodes Including initial geographic location, power status and data collection priority, sensor nodes Divided into clusters, each cluster contains a group of sensor nodes, The sensor nodes in a cluster are represented as: HP k SN = {SN k1 , SN k2 , ..., SN kl}, , the center point of the cluster is the initial hovering point HP, denoted as: H1, H2, ..., H m ;

[0015] Step 2.2, calculate the initial weight: according to the sensor node The weight of the sensor node is calculated based on the initial geographical location, power status and data collection priority :

[0016]

[0017] in, For data collection priority, The power state of the sensor node initialization, and is the weight adjustment coefficient, For the sensor nodes;

[0018] Step 2.3: According to the weight of sensor nodes , use the weighted K-Medoids algorithm to perform initial node clustering and minimize the weighted clustering loss function :

[0019]

[0020] in: For sensor nodes and hover point The Euclidean distance, let SNi The coordinates of (x i ,y i ), H k The coordinates of (x k ,y k ),but , For all the nodes in the current k-th cluster that are assigned to the hovering point H k The node set of

[0021] Step 2.4, determine the position of the initial hovering point HP: use the weighted K-Medoids algorithm to determine the center point of each cluster as the initial hovering point HP of the drone;

[0022] Step 2.5, Coverage and Overlap Judgment: The coverage range of each initial hovering point HP is set with the coverage radius R of the given drone, and the overlapping area between the initial hovering points HP is identified. For the overlapping sensor nodes, the drone wireless charging network calculates the attribution score based on the node attributes of the overlapping sensor nodes and uses the following attribution judgment function to assign them:

[0023]

[0024] in, Represents a sensor node Home to the initial hover point Rating, For all the nodes in the current k-th cluster that are assigned to the hovering point H k The node set of is the regulating factor, Represents a sensor node The current remaining power, Represents a sensor node The data collection priority, Represents a node collection The total power of all nodes in Represents a node collection The data collection priority of all nodes in the sum is used to assign each overlapping node to the initial hovering point HP with the highest score, ensuring that each sensor node belongs to only one initial hovering point HP to prevent multiple scheduling or task conflicts;

[0025] Step 2.6: Output the initial hovering point (HP) plan: Finally determine the initial hovering point layout and node division results of the UAV.

[0026] A further improvement of the present invention is that in step 3, dynamically optimizing the data collection path, the initial hovering point (HP) access sequence, and the hovering time of the UAV specifically includes the following steps:

[0027] Step 3.1: Establish a dual deep Q network (DDQN) reinforcement learning algorithm and determine the dual depth The state space, action space, and reward function of the network (DDQN) reinforcement learning algorithm;

[0028] Step 3.2: Based on the state space, action space and reward function in step 3.1, use the dual depth The network (DDQN) reinforcement learning algorithm plans the flight path of the UAV in real time and dynamically determines the UAV's stay time and charging task at each initial hovering point (HP).

[0029] A further improvement of the present invention is that in step 3.1, the state space includes the current position of the UAV, the power of the sensor nodes, the data collection priority, the age of information (AoI) of each sensor node, the remaining energy of the UAV, and the remaining time for task execution. Specifically, the state space vector is expressed as:

[0030] in, is the current position of the drone, is the power status of the sensor node, is the data collection priority of the sensor node, is the current age of information (AoI) of each sensor node, is the remaining energy of the drone, The remaining time for task execution.

[0031] A further improvement of the present invention is that in step 3.1, the action space UAV selects the next hovering point (HP), the size of the action space is the number of hovering points (HP) M, and the UAV selects the next hovering point position from the current hovering point position. Each action represents that the UAV flies from the current hovering point position to the next hovering point position, and determines the stay time at the point to perform charging and data collection tasks:

[0032] ,in, is the set of action spaces.

[0033] A further improvement of the present invention is that in step 3.1, the reward function aims to optimize the age of information (AoI) of the system:

[0034] in: is the average age of information (AoI) within the cluster after the drone completes data collection at the current hovering point, is the average age of information (AoI) of the wireless rechargeable sensor network after the current task is executed, To adjust the weight parameter of the optimized ratio of age of information (AoI) within the cluster, Weight parameter for adjusting the global Age of Information (AoI) optimization ratio.

[0035] A further improvement of the present invention is that step 4 specifically includes the following steps:

[0036] Step 4.1: After each charging mission is completed, the drone returns to the data center and feeds back all sensor node information of this mission to the dual depth Network (DDQN) reinforcement learning algorithm;

[0037] Step 4.2: Double Depth The network (DDQN) reinforcement learning algorithm dynamically adjusts the clustering weights of sensor nodes based on the real-time feedback of sensor node information, so that the clustering weights are dynamically associated with the age of information (AoI) status of the sensor nodes to optimize the HP selection of the next task. The specific weight adjustment method is:

[0038]

[0039] in: is the current Age of Information (AoI) value of the node, Double depth The parameters of the network (DDQN) reinforcement learning algorithm are automatically adjusted according to the feedback results. The higher the Age of Information (AoI) value of the sensor node, the greater the weight in the next round of clustering, and the easier it is to be optimized and assigned to the appropriate HP.

[0040] Step 4.3: Utilize Double Depth The adjusted weights output by the network (DDQN) reinforcement learning algorithm are re-applied to the weighted K-Medoids clustering algorithm, and steps 2.3 to 2.5 are performed again to obtain a new division of hovering points HP and sensor nodes SN;

[0041] Step 4.4, Step 4.3 updated hover point layout and double depth The network (DDQN) was again applied to the next drone data collection mission.

[0042] The beneficial effects of this application are:

[0043] This application uses the real-time feedback mechanism of the DDQN algorithm to quickly respond to changes in the network environment and dynamically adjust the drone path and hovering strategy, significantly improving the execution efficiency and flexibility of data collection tasks.

[0044] This application can adapt to network node changes such as node location changes, new node additions, or existing node failures, and automatically adjust clustering and path planning strategies. The system can maintain a high level of adaptability without human intervention.

[0045] The reinforcement learning and clustering algorithms used in the system of this application have good modular characteristics. When the network scale or task requirements expand, this application can be directly expanded by adding drones or adjusting the number of clusters, with almost no additional increase in system complexity or operating costs, and has strong scalability.

[0046] To sum up, this application can effectively reduce data collection delays, improve the real-time nature of information collection and the overall network performance. It has the advantages of high efficiency, flexibility, strong adaptability and scalability, and can be widely used in fields such as Internet of Things monitoring, smart agriculture and environmental monitoring. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 This is the network architecture diagram of this application.

[0048] Figure 2 is a flowchart of the initial hover point (HP) selection for this application.

[0049] Figure 3 This is the flow chart of the DDQN algorithm of this application.

[0050] Figure 4 This is a comparison chart of the iterative results of this application and other algorithms.

[0051] Figure 5 This is a comparison chart of the acquisition success rate of this application and other algorithms. DETAILED DESCRIPTION

[0052] The following drawings illustrate embodiments of the present invention. For clarity, many practical details will be included in the following description. However, it should be understood that these practical details are not intended to limit the present invention. In other words, in some embodiments of the present invention, these practical details are not essential. Furthermore, to simplify the drawings, some commonly used structures and components are depicted in a simplified schematic manner.

[0053] like Figure 1 As shown, the present application is a method for optimizing the age of data collection information in a wireless rechargeable sensor network, and the method specifically includes the following steps:

[0054] Step 1: Build the network architecture. The network architecture is a wireless drone charging network (WRSN), consisting of multiple sensor nodes (SNs) and a data center (DC). The DC serves as the core node of the network, responsible for task scheduling, path planning, and data processing. Drones depart from the DC to collect data and charge the sensor nodes (SNs). Upon completion, they return to the DC to upload the collected data for further processing.

[0055] Hovering points (HPs) are key locations for drones to perform missions. Located in the core area of ​​sensor node distribution, they cover multiple sensor nodes within a certain range. HPs are selected using a weighted K-Medoids clustering algorithm, which divides sensor nodes based on their geographic location and mission requirements. The center of each cluster is designated as the HP, where the drone rests, charges, and collects data.

[0056] During the mission, the drone charges the SN via WPT technology and exchanges data with the SN via short-range communication, storing the collected monitoring data in real time on its own device. Subsequently, the DC, based on the deep reinforcement learning (DDQN) algorithm, dynamically adjusts the drone's hovering position, access sequence, and charging strategy, taking into account the battery status, data priority, and geographic distribution of sensor nodes to optimize the Age of Information (AoI), thereby improving the real-time performance and overall performance of the entire network.

[0057] Step 2: Initial hovering point (HP) selection.

[0058] In this application, the selection of HP (hover point) is one of the key decisions when the UAV performs data collection and charging tasks. By reasonably selecting the HP location, the data collection efficiency can be effectively improved and the age of information (AoI) can be reduced. AoI is defined as the time interval from the generation of the latest information by the sensor node (SN) to the completion of the collection by the UAV. By optimizing the location of HP and the division of clusters, the average AoI of the SN nodes in each cluster is minimized, which is one of the main goals of the HP selection stage. Specifically, this application adopts a weighted K-Medoids clustering algorithm to divide the sensor nodes in the UAV wireless charging network (WRSN) into multiple clusters, and uses the center point of each cluster as the initial hovering point (HP) of the UAV, and finally determines the initial hovering point layout and node division results of the UAV;

[0059] The initial hover point (HP) selection process includes the following steps:

[0060] Step 2.1, Initialize network parameters: Set the coverage radius of the drone to , there are a total of sensor nodes (SN), each of which includes an initial geographical location, power status, and data collection priority. Sensor nodes (SN) are divided into clusters, each cluster contains a group of relatively concentrated sensor nodes. The sensor nodes in a cluster are represented as: HP k SN = {SN k1 , SN k2 , ..., SN kl}, , the center point of the cluster is the initial hovering point (HP), denoted as: H1, H2, ..., H m ;

[0061] Step 2.2, calculate the initial weight: calculate the weight of the sensor node according to the initial geographical location, power status and data collection priority of the sensor node (SN) :

[0062]

[0063] in: For data collection priority, The power status of the sensor node when it is initialized. 0 means no power at all. and is the weight adjustment coefficient;

[0064] Step 2.3: According to the weight of sensor nodes , use the weighted K-Medoids algorithm to perform initial node clustering and minimize the weighted clustering loss function :

[0065]

[0066] in: For sensor nodes and hover point The Euclidean distance, let SN i The coordinates of (x i ,y i ), H k The coordinates of (x k ,y k ),but , For all the nodes in the current k-th cluster that are assigned to the hovering point H k The node set of

[0067] Step 2.4, determine the position of the initial hovering point HP: use the weighted K-Medoids algorithm to determine the center point of each cluster as the initial hovering point HP of the drone;

[0068] Step 2.5, Coverage and Overlap Judgment: The coverage range of each initial hovering point HP is set with the coverage radius R of the given drone, and the overlapping area between the initial hovering points HP is identified. For the overlapping sensor nodes, the drone wireless charging network calculates the attribution score based on the node attributes of the overlapping sensor nodes and uses the following attribution judgment function to assign them:

[0069]

[0070] in, Represents a sensor node Home to the initial hover point Rating, For all the nodes in the current k-th cluster that are assigned to the hovering point H k The node set of is the regulating factor, Represents a sensor node The current remaining power, Represents a sensor node The data collection priority, Represents a node collection The total power of all nodes in Represents a node collection The data collection priority of all nodes in the sum is used to assign each overlapping node to the initial hovering point HP with the highest score, ensuring that each sensor node belongs to only one initial hovering point HP to prevent multiple scheduling or task conflicts;

[0071] Step 2.6: Output the initial hovering point (HP) plan: Finally determine the initial hovering point layout and node division results of the UAV.

[0072] The aforementioned objective function of initial node clustering ensures that the selected HPs more appropriately cover clusters of nodes of high importance (e.g., low battery, high data priority), thereby improving initial data collection efficiency and charging effectiveness, and thereby reducing the initial node AoI. Furthermore, in actual implementation, due to the potential overlap of drone coverage areas (coverage radius R), some SN nodes may be covered by multiple clusters simultaneously. To avoid overlapping coverage or conflicts, the present invention utilizes a specialized decision algorithm to adjust node ownership, ensuring that each SN node belongs to only one HP.

[0073] The HP layout plan and cluster division results output in this stage will serve as the initial input for the subsequent DDQN reinforcement learning path planning stage, and the reinforcement learning feedback mechanism will further achieve dynamic optimization of the AoI within the cluster and the overall system.

[0074] Step 3: Using the position of the initial hovering point (HP) of the drone determined in step 2 and the node division result as the initial layout, use the double depth A DDQN reinforcement learning algorithm dynamically optimizes the drone's data collection path, initial hovering point (HP) visit sequence, and hovering time, enabling the drone to collect data and charge each cluster node. Through the DDQN reinforcement learning feedback mechanism, the clustering weights and path planning strategies selected by the HP are continuously adjusted to minimize the overall system average AoI.

[0075] AoI (Age of Information) is defined as the time from when the sensor node (SN) generates the latest data to when the drone successfully receives the data. The overall goal of the system is to minimize the average AoI of all nodes:

[0076]

[0077] The implementation steps of the DDQN algorithm are as follows Figure 3 As shown in Figure 2, DDQN is a reinforcement learning method widely used in the field of deep reinforcement learning. Compared with the traditional DQN method, DDQN has higher stability and convergence speed. The core idea of ​​DDQN is to use two independent Network, in which a main network (Main Network) is used to select actions and a target network (Target Network) is used to evaluate actions, thus avoiding the possible The problem of overestimation of value. The state-action value function in the DDQN algorithm ( Function) update formula is as follows:

[0078]

[0079] in: is the state at time t, describing the current network state including drone location, SN battery status, AoI status, etc. is the action at time t, indicating the selected next visited hover point (HP); Immediate rewards for reinforcement learning; is the learning rate; is the discount factor; For the target network value;

[0080] The state space in this application includes the current position of the drone, the power of the sensor nodes, the data collection priority, the age of information (AoI) of each sensor node, the remaining energy of the drone and the remaining time for mission execution. Specifically, the state space vector is represented as:

[0081] in, is the current position of the drone, is the power status of the sensor node, is the data collection priority of the sensor node, is the current age of information (AoI) of each sensor node, is the remaining energy of the drone, The remaining time for task execution.

[0082] The action space is used by the drone to select the next hovering point (HP). The size of the action space is the number of hovering points (HP), M, where M is the number of clusters, and each cluster has one hovering point HP. The drone selects the next hovering point from the current hovering point. Each action indicates that the drone flies from the current hovering point to the next hovering point, and determines the dwell time at the point to perform charging and data collection tasks:

[0083] ,

[0084] in, is the set of action spaces.

[0085] The reward function aims to optimize the system's Age of Information (AoI):

[0086] in: is the average age of information (AoI) within the cluster after the drone completes data collection at the current hovering point, is the average age of information (AoI) of the wireless rechargeable sensor network after the current task is executed, To adjust the weight parameter of the optimized ratio of age of information (AoI) within the cluster, Weight parameter for adjusting the global Age of Information (AoI) optimization ratio.

[0087] Through this reward function design, the reinforcement learning algorithm will strive to reduce the AoI at both the local (within the cluster) and global levels, achieve overall optimization, and minimize the system AoI.

[0088] Step 4: After each charging task is executed, according to the double depth The feedback results of the network (DDQN) reinforcement learning algorithm dynamically adjust the weights of the nodes in the weighted K-Medoids clustering algorithm, so that the weights of the sensor nodes in the weighted K-Medoids clustering algorithm are proportional to the current age of information (AoI) of the sensor nodes, achieving dynamic optimization of the average age of information (AoI) within the cluster and the overall average age of information (AoI) of the wireless charging network (WRSN). Step 4 specifically includes the following steps:

[0089] Step 4.1: After each charging task is completed, the drone returns to the data center and feeds back all the sensor node information of this task, i.e. data update status, power change, and AoI status, to the dual depth Network (DDQN) reinforcement learning algorithm;

[0090] Step 4.2: Double Depth The network (DDQN) reinforcement learning algorithm dynamically adjusts the clustering weights of sensor nodes based on the real-time feedback of sensor node information, so that the clustering weights are dynamically associated with the age of information (AoI) status of the sensor nodes to optimize the HP selection of the next task. The specific weight adjustment method is:

[0091]

[0092] in: is the current Age of Information (AoI) value of the node, Double depth The parameters of the network (DDQN) reinforcement learning algorithm are automatically adjusted according to the feedback results. The higher the Age of Information (AoI) value of the sensor node, the greater the weight in the next round of clustering, and the easier it is to be optimized and assigned to the appropriate HP.

[0093] Step 4.3: Utilize Double Depth The adjusted weights output by the reinforcement learning algorithm of the network (DDQN) are re-performed with the weighted K-Medoids clustering algorithm, and steps 2.3 to 2.5 are executed again to obtain a new division of hovering points HP and sensor nodes SN to optimize the overall path planning and information collection efficiency.

[0094] Step 4.4, Step 4.3 updated hover point layout and double depth The network (DDQN) was again applied to the next drone data collection mission.

[0095] Step 5: Repeat steps 3 and 4 to continue the drone data collection and charging tasks until the average age of information (AoI) of the drone wireless charging network (WRSN) reaches a stable optimization state or reaches the preset maximum number of optimization iterations.

[0096] like Figure 4 The comparison of the iterative results of this application and other algorithms is shown. Figure 4 It can be seen that the algorithm (KM-DDQN) of this application converges faster and has a smaller average AoI than the K-Medoids shortest path algorithm (KM-SPP) and the traditional deep Q network algorithm (DQN).

[0097] Figure 5 The comparison of the data collection success rate of this application and other algorithms is shown. Figure 5 It can be seen that the algorithm of this application (KM-DDQN) has a higher data collection success rate than the K-Medoids shortest path algorithm (KM-SPP) and the traditional deep Q network algorithm (DQN) in an environment with the same number of sensor nodes.

[0098] This application proposes a drone path planning method that combines a weighted K-Medoids clustering algorithm with DDQN reinforcement learning, demonstrating significant advantages in wireless charging networks (WRSNs). To clearly analyze and evaluate the technical performance of this invention, the following will provide a detailed analysis of the system's timeliness, rationality, fairness, and reliability.

[0099] 1. Timeliness.

[0100] This application aims to minimize the overall age of information (AoI) of the network. It uses the DDQN algorithm to dynamically optimize drone paths and hovering methods, making data collection more real-time. Compared with traditional static path planning, this application can more effectively reduce data collection latency, significantly reducing the average AoI of sensor node information, and thus significantly improving the timeliness of network data updates.

[0101] 2. Reasonableness.

[0102] This application uses a weighted K-Medoids clustering algorithm to rationally select hovering points (HPs) in the initial phase, while explicitly considering factors such as node data priority and battery status to ensure more scientific and reasonable node allocation. Furthermore, the DDQN algorithm's dynamic feedback adjustment mechanism continuously optimizes HP selection, further improving the rationality of system operation.

[0103] 3. Fairness.

[0104] This application explicitly incorporates a node weight adjustment mechanism into the clustering and reinforcement learning process, giving higher optimization priority to nodes with low battery and urgent information update needs, thereby achieving fair distribution of system resources. Through continuous iterative optimization, the system can gradually eliminate the imbalance in resource allocation and ensure fair and effective data update opportunities for each SN node.

[0105] 4. Reliability.

[0106] This application introduces the adaptive characteristics of DDQN reinforcement learning, which can perceive and respond to changes in the network environment in real time, such as insufficient node power, node failure, or demand changes. The drone can adjust its strategy in time based on real-time feedback, significantly improving the robustness and reliability of the system operation.

[0107] The foregoing is merely an embodiment of the present invention and is not intended to limit the present invention. It will be apparent to those skilled in the art that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention are intended to be included within the scope of the claims of the present invention.

Claims

1. A method for optimizing the age of data collection information in a wireless rechargeable sensor network, characterized by: The method for optimizing the age of data collection information in a wireless rechargeable sensor network specifically comprises the following steps: Step 1: Build a network architecture: The network architecture is a UAV wireless charging network, which includes multiple sensor nodes and a data center. UAVs depart from the data center to perform data collection and wireless charging tasks, and return to the data center after the mission is completed. Step 2: Initial hovering point selection: Use the weighted K-Medoids clustering algorithm to divide the sensor nodes in the UAV wireless charging network into multiple clusters, and use the center point of each cluster as the initial hovering point of the UAV to ultimately determine the hovering point layout and node division results of the UAV; Step 3: Using the layout of the drone's hovering points and the node division results determined in step 2 as the initial layout, use the double depth The network reinforcement learning algorithm dynamically optimizes the drone's data collection path, initial hovering point visit sequence, and hovering time, enabling the drone to collect data and charge each cluster node. Step 4: After each charging task is executed, according to the double depth The feedback results of the network reinforcement learning algorithm dynamically adjust the weights of the nodes in the weighted K-Medoids clustering algorithm, making the weights of the sensor nodes in the weighted K-Medoids clustering algorithm proportional to the current information age of the sensor nodes, thereby achieving dynamic optimization of the average information age within the cluster and the overall average information age of the drone wireless charging network; Step 5: Repeat steps 3 and 4 to continue the drone data collection and charging tasks until the average information age of the drone wireless charging network reaches a stable optimization state or reaches a preset maximum number of optimization iterations. Step 2 uses a weighted K-Medoids clustering algorithm to determine the initial hovering point layout and node division results of the drone, which specifically includes the following steps: Step 2.1, Initialize network parameters: Set the coverage radius of the drone to , there are a total of sensor nodes , each of the sensor nodes Including initial geographic location, power status and data collection priority, sensor nodes Divided into clusters, each cluster contains a group of sensor nodes, The sensor nodes in a cluster are represented as: , , the center point of the cluster is the initial hovering point HP, recorded as: ; Step 2.2, calculate the initial weight: according to the sensor node The weight of the sensor node is calculated based on the initial geographical location, power status and data collection priority : , in, For data collection priority, The power state of the sensor node initialization, and is the weight adjustment coefficient, For the sensor nodes; Step 2.3: According to the weight of sensor nodes , use the weighted K-Medoids algorithm to perform initial node clustering and minimize the weighted clustering loss function : , in: For sensor nodes and hover point The Euclidean distance of For all the hovering points H in the current k-th cluster that are assigned to the initial k The node set of Step 2.4, determine the initial hovering point HP position: use the weighted K-Medoids algorithm to determine the center point of each cluster as the initial hovering point HP of the drone; Step 2.5, Coverage and Overlap Judgment: The coverage range of each initial hovering point HP is set with the coverage radius R of the given drone, and the overlapping area between the initial hovering points HP is identified. For the overlapping sensor nodes, the drone wireless charging network calculates the attribution score based on the node attributes of the overlapping sensor nodes and uses the following attribution judgment function to assign them: , in, Represents a sensor node Home to the initial hover point Rating, For all the hovering points in the current k-th cluster that are assigned to the initial The node set of is the regulating factor, Represents a sensor node The current remaining power, Represents a sensor node The data collection priority, Represents a node collection The total power of all nodes in Represents a node collection The data collection priority of all nodes in the sum is used to assign each overlapping node to the initial hovering point HP with the highest score, ensuring that each sensor node belongs to only one initial hovering point HP to prevent multiple scheduling or task conflicts; Step 2.6: Output the hovering point plan: Finally determine the hovering point layout and node division results of the drone.

2. The method for optimizing the age of data collection information in a wireless rechargeable sensor network according to claim 1, characterized in that: In step 3, the dynamic optimization of the drone's data collection path, initial hovering point visit sequence, and hovering time specifically includes the following steps: Step 3.

1. Create double depth Network reinforcement learning algorithm to determine double depth The state space, action space, and reward function of network reinforcement learning algorithms; Step 3.2: Based on the state space, action space and reward function in step 3.1, use the dual depth The network reinforcement learning algorithm plans the UAV's flight path in real time and dynamically determines the UAV's stay time and charging tasks at each initial hovering point.

3. The method for optimizing the age of data collection information in a wireless rechargeable sensor network according to claim 2, characterized in that: In step 3.1, the state space includes the current position of the UAV, the power level of the sensor nodes, the data collection priority, the information age of each sensor node, the remaining energy of the UAV, and the remaining time for mission execution. Specifically, the state space vector is expressed as: , in, is the current position of the drone, is the power status of the sensor node, is the data collection priority of the sensor node, is the current information age of each sensor node, is the remaining energy of the drone, The remaining time for task execution.

4. The method for optimizing the age of data collection information in a wireless rechargeable sensor network according to claim 2, characterized in that: In step 3.1, the action space is used by the drone to select the next hovering point. The size of the action space is the number of hovering points M. The drone selects the next hovering point from the current hovering point. Each action indicates that the drone flies from the current hovering point to the next hovering point and decides the stay time at the point to perform charging and data collection tasks: , in, is the set of action spaces.

5. The method for optimizing the age of data collection information in a wireless rechargeable sensor network according to claim 4, characterized in that: In step 3.1, the reward function aims to optimize the information age of the system: , in: is the average information age within the cluster after the UAV completes data collection at the current hovering point, is the average information age of the wireless rechargeable sensor network after the current task is executed, To adjust the weight parameter of the optimal ratio of information age within the cluster, The weight parameter for adjusting the global information age optimization ratio.

6. The method for optimizing the age of data collection information in a wireless rechargeable sensor network according to claim 5, characterized in that: Step 4 specifically includes the following steps: Step 4.1: After each charging mission is completed, the drone returns to the data center and feeds back all sensor node information of this mission to the dual depth Network reinforcement learning algorithm; Step 4.2: Double Depth The network reinforcement learning algorithm dynamically adjusts the clustering weight of the sensor nodes based on the real-time feedback of the sensor node information, so that the clustering weight is dynamically associated with the information age AoI state of the sensor nodes to optimize the HP selection of the next task. The specific weight adjustment method is: , in: For sensor nodes The current information age value, Double depth The network (DDQN) reinforcement learning algorithm automatically adjusts the parameters according to the feedback results. The higher the information age value of the sensor node, the greater the weight in the next round of clustering, and the easier it is to be optimized and assigned to the appropriate HP; Step 4.3: Utilize Double Depth The adjusted weights output by the network reinforcement learning algorithm are re-applied to the weighted K-Medoids clustering algorithm, and steps 2.3 to 2.5 are executed again to obtain a new division of hovering points HP and sensor nodes SN. Step 4.4, Step 4.3 updated hover point layout and double depth The network was used again in the next drone data collection mission.

Citation Information

Patent Citations

  • Unmanned aerial vehicle track planning and video transmission method based on cellular network

    CN117768987A

  • Space-air-ground integrated UAV-assisted IoT data collectioncollection method based on aoi

    US20230239037A1