AoI-driven AUV-assisted underwater Internet of Things high-energy-efficiency data acquisition method
By constructing a spatiotemporally decoupled hierarchical collaborative architecture and multi-AUV collaborative scheduling, the problems of energy consumption, timeliness and network stability in underwater IoT data acquisition are solved, and efficient and reliable data acquisition is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-26
- Publication Date
- 2026-04-14
AI Technical Summary
Existing AUV-assisted underwater IoT data acquisition solutions struggle to meet dynamic upload demands under multiple constraints in complex underwater environments, particularly issues related to equipment status changes and data timeliness. This leads to network failures and redundant energy consumption, and traditional multi-AUV scheduling lacks online adaptability.
A spatiotemporally decoupled hierarchical collaborative architecture is constructed, with USV as the high-level decision-maker and AUV as the low-level decision-maker. By combining communication models, data acquisition models and AoSI models, Markov decision processes and HRL-GIC hierarchical reinforcement learning algorithms are used to optimize task allocation and execution, thereby achieving multi-AUV collaborative scheduling.
It optimizes energy consumption, extends device battery life, ensures data timeliness and network stability, adapts to complex environments, and meets the dynamic upload needs of large-scale IoUT devices.
Smart Images

Figure CN121865299A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this application relate to the fields of underwater Internet of Things (IoT) and mobile swarm intelligence sensing technology, and particularly to an AoI-driven AUV-assisted high-efficiency data acquisition method for underwater IoT. Background Technology
[0002] The Internet of Underwater Things (IoUT) is an intelligent network comprised of various interconnected underwater sensing devices. IoUTs utilize these devices and related technologies to sense, analyze, and transmit data related to the underwater environment, providing technical support for underwater search and resource exploration activities. Unlike terrestrial wireless communication environments, the underwater environment is complex and variable, and electromagnetic signals experience severe attenuation underwater, rendering radio frequency communication technologies unsuitable for IoUTs. Underwater acoustic communication technology is widely used as an alternative, but IoUT networks based on underwater acoustic communication suffer from low bandwidth and high latency, making them unsuitable for long-distance transmission of large amounts of data. Therefore, designing an efficient and reliable data collection scheme within the IoUT network remains a core research challenge.
[0003] Traditional data collection methods typically utilize multi-hop routing between IoUT devices, which increases energy consumption and consequently impacts the lifespan of battery-powered devices without recharging capabilities. To address this challenge, hierarchical data collection schemes have been extensively explored. In these schemes, devices are first divided into multiple clusters, with designated cluster heads responsible for collecting data from other devices within their respective clusters. These cluster heads then transmit the collected data to a surface base station. However, these cluster heads often fail prematurely due to increased communication overhead, potentially causing network failures. In contrast, deploying AUVs (Autonomous Underwater Vehicles) as mobile platforms to collect data via underwater acoustic links has proven to be a more cost-effective and efficient solution.
[0004] Current AUV-assisted data collection schemes typically rely on preset trajectories, requiring the equipment to forward data to the AUV along its preset path. This design inevitably leads to redundant energy consumption.
[0005] Many studies have been conducted to address this issue, but most of these studies focus on optimizing AUV trajectories by minimizing travel distance. This approach ignores the impact of TE (Turbulent Environment), communication limitations, and device upload requirements on the trajectory.
[0006] The real-time status of IoUT devices varies depending on their geographical location and operational tasks, but research on this variability is limited. Specifically, devices with large backlogs of stored data and high environmental data generation rates need to be prioritized for access by AUVs, because if data is not collected in a timely manner, it will lead to buffer overflow, with old data being overwritten by new data. More importantly, in time-sensitive tasks such as underwater environmental monitoring and reconnaissance, the value of data decays rapidly over time. Simply pursuing data throughput while ignoring the AoI (Age of Information) of the data may result in the system collecting a large amount of outdated and invalid information. Under these multiple constraints, a single AUV obviously cannot meet the dynamic upload requirements of a large number of IoUT devices in a timely manner, thus necessitating research on collaborative scheduling schemes for multiple AUVs.
[0007] However, current commonly used multi-AUV scheduling schemes are usually designed for specific static tasks. The traditional optimization methods used often rely on a large amount of prior environmental data and lack online adaptability, making it difficult to cope with high-dimensional multi-objective optimization problems in complex underwater environments. Summary of the Invention
[0008] To address the aforementioned technical issues, embodiments of this application propose an AoI-driven AUV-assisted underwater IoT high-efficiency data acquisition method. This method constructs a spatiotemporally decoupled hierarchical collaborative architecture, achieving efficient data acquisition in complex underwater environments through the organic combination of high-level task allocation and low-level distributed execution, thereby maximizing the data acquisition ratio and average sensing coverage.
[0009] To achieve the above objectives, embodiments of this application propose an AoI-driven AUV-assisted high-efficiency data acquisition method for underwater IoT, implemented based on a spatiotemporally decoupled hierarchical collaborative architecture. The method includes: constructing a spatiotemporally decoupled hierarchical collaborative architecture composed of multiple agents for underwater data acquisition scenarios. In this architecture, the USV acts as the high-level decision-maker, responsible for high-level task allocation, while each AUV acts as the low-level decision-maker, responsible for low-level task execution; and constructing a communication model, a data acquisition model, and an AoSI model. The communication model represents the underwater acoustic communication between each IoUT device and each AUV, and the communication between each AUV and the USV. Underwater acoustic communication and radio frequency communication between USVs and shore-based base stations are described. The data acquisition model represents the data acquisition by each AUV from each IoUT device, and the AoSI model is used to measure the timeliness and semantic value of data acquisition. An optimization problem is defined based on the hierarchical collaborative architecture, communication model, data acquisition model, and AoSI model, and a Markov decision process is used to model the optimization problem. The modeling results of the optimization problem are processed using the HRL-GIC hierarchical reinforcement learning algorithm to obtain the global optimal policy. This policy instructs the hierarchical collaborative architecture to perform underwater data acquisition based on the optimal task allocation scheme and optimal task execution scheme in the global optimal policy.
[0010] To achieve the above objectives, embodiments of this application also propose an AoI-driven AUV-assisted underwater IoT high-efficiency data acquisition system, implemented based on the AoI-driven AUV-assisted underwater IoT high-efficiency data acquisition method described above. The system includes: a hierarchical collaborative architecture construction module, used to construct a spatiotemporally decoupled hierarchical collaborative architecture composed of multiple agents for underwater data acquisition scenarios. In this hierarchical collaborative architecture, the USV acts as the high-level decision-maker, responsible for high-level task allocation, and each AUV acts as the low-level decision-maker, responsible for low-level task execution; and a model construction module, used to construct a communication model, a data acquisition model, and an AoSI model. The communication model represents the underwater acoustic communication between each IoUT device and each AUV, and the underwater acoustic communication between each AU... The underwater acoustic communication between AUVs and USVs, and the radio frequency communication between USVs and shore-based base stations, are described. The data acquisition model represents the data acquisition by each AUV from its respective IoUT devices. The AoSI model is used to measure the timeliness and semantic value of data acquisition. The optimization problem modeling module is used to define the optimization problem based on the hierarchical collaborative architecture, communication model, data acquisition model, and AoSI model, and to model the optimization problem using a Markov decision process. The optimization search module is used to process the modeling results of the optimization problem using the HRL-GIC hierarchical reinforcement learning algorithm to obtain the globally optimal strategy. The execution module is used to instruct the hierarchical collaborative architecture to perform underwater data acquisition based on the optimal task allocation scheme and optimal task execution scheme in the globally optimal strategy.
[0011] To achieve the above objectives, embodiments of this application also propose an electronic device, including a processor and a memory, wherein the memory stores instructions executable by the processor, and the processor is configured to execute the instructions such that the electronic device can implement an AoI-driven AUV-assisted underwater IoT high-efficiency data acquisition method as described above.
[0012] To achieve the above objectives, embodiments of this application also propose a computer-readable storage medium storing a computer program that, when executed by a processor, enables the implementation of an AoI-driven AUV-assisted underwater IoT high-efficiency data acquisition method as described above.
[0013] The embodiments of this application propose an AoI-driven AUV-assisted high-efficiency data acquisition method for underwater IoT. Through innovative architecture design and algorithm optimization, it specifically addresses multiple core pain points faced by underwater IoT data acquisition, bringing significant benefits in multiple dimensions.
[0014] In terms of energy consumption optimization and device battery life, the solution adopts a spatiotemporally decoupled layered collaborative architecture that coordinates task allocation by USV and performs data collection by AUV, replacing the traditional multi-hop routing and fixed cluster head mode. This not only avoids the extra energy consumption caused by multi-hop transmission, but also reduces the redundant energy consumption caused by AUV preset trajectories through global optimal task allocation, effectively extending the operating life of IoUT devices powered by batteries without charging capabilities.
[0015] In terms of data timeliness and value assurance, the specially constructed AoSI model can accurately measure the timeliness and semantic value of data collection. Combined with the multi-AUV collaborative scheduling strategy, it can prioritize access (collection) of devices with large backlogs of stored data and high environmental data generation rates, thereby avoiding the problem of old data being overwritten due to buffer overflows. At the same time, for time-sensitive tasks such as underwater environmental monitoring and reconnaissance, it minimizes AoI and eliminates the collection of outdated and invalid information, ensuring the actual application value of the data.
[0016] In terms of adaptability to complex environments and scheduling efficiency, this application integrates communication models of underwater acoustic communication and radio frequency communication, which fully adapts to the communication characteristics of complex underwater environments and makes up for the shortcomings of low bandwidth and long latency of single underwater acoustic communication. It models the optimization problem through Markov decision process and uses the HRL-GIC hierarchical reinforcement learning algorithm to solve the global optimal strategy. It gets rid of the dependence of traditional multi-AUV scheduling on a large amount of prior environmental data, has strong online adaptability, and can efficiently cope with high-dimensional multi-objective optimization problems under multiple constraints such as turbulent environment, communication limitations, and dynamic changes in equipment status, and meet the dynamic upload requirements of large-scale IoUT devices.
[0017] In terms of network stability, replacing the traditional cluster head centralized collection mode with AUV mobile acquisition completely solves the problem of premature failure of the cluster head due to excessive communication overhead, significantly reduces the risk of potential network failures, and ensures the continuous reliability of the IoUT data acquisition process. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in the embodiments or related technologies of this application, the accompanying drawings used in the description of the embodiments or related technologies of this application will be briefly introduced below. Obviously, the following drawings are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. The drawings described herein are only used to explain this application and are not intended to limit this application.
[0019] Figure 1 This is a flowchart of an AoI-driven AUV-assisted underwater IoT high-efficiency data acquisition method provided in one embodiment of this application; Figure 2 This is a schematic diagram of an underwater data acquisition scenario provided in one embodiment of this application; Figure 3 This is a schematic diagram of the HRL-GIC hierarchical reinforcement learning algorithm provided in one embodiment of this application; Figure 4 This is a schematic diagram of a high-efficiency data acquisition structure for AoI-driven AUV-assisted underwater IoT provided in another embodiment of this application; Figure 5 This is a schematic diagram of the structure of an electronic device provided in another embodiment of this application. Detailed Implementation
[0020] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the various embodiments of this application will be described in detail below with reference to the accompanying drawings. Those skilled in the art will understand that many technical details have been presented in the embodiments of this application to facilitate better understanding. However, the technical solutions claimed in this application can be implemented even without these technical details and various variations and modifications based on the following embodiments. The division of the following embodiments is for ease of description and should not constitute any limitation on the specific implementation of this application. The following embodiments can be combined with and referenced by each other without contradiction.
[0021] One embodiment of this application proposes an AoI-driven AUV-assisted high-efficiency data acquisition method for underwater IoT, which is implemented based on a spatiotemporally decoupled hierarchical collaborative architecture. The implementation details of the AoI-driven AUV-assisted high-efficiency data acquisition method for underwater IoT proposed in this embodiment are described below. The following content is only for the convenience of understanding the implementation details and is not necessary for implementing this solution.
[0022] The specific process of the AoI-driven AUV-assisted underwater IoT high-efficiency data acquisition method proposed in this embodiment can be described as follows: Figure 1 As shown, it includes: Step 11: For underwater data acquisition scenarios, construct a spatiotemporally decoupled hierarchical collaborative architecture composed of multiple agents. In the hierarchical collaborative architecture, the USV acts as the high-level decision-maker, responsible for high-level task allocation, while each AUV acts as the low-level decision-maker, responsible for low-level task execution.
[0023] In practical implementation, the first step in data acquisition is to construct a spatiotemporally decoupled hierarchical collaborative architecture composed of multiple agents, specifically for underwater data acquisition scenarios. In this hierarchical collaborative architecture, the USV acts as the high-level decision-maker, responsible for high-level task allocation, while each AUV acts as the low-level decision-maker, responsible for low-level task execution.
[0024] In one example, an underwater data acquisition scenario is as follows: Figure 2 As shown, in an underwater data acquisition scenario, multiple AUVs continuously cruise in a turbulent ocean environment and, based on the status information of the IoUT devices broadcast by the USVs, determine the target IoUT devices that need to be accessed for data acquisition.
[0025] Assume there is a shore-based base station and a USV. Taiwan AUV and Taiwan IoUT device, Taiwan AUV is a collection express, Indicates the first Taiwanese AUV, A single IoUT device consists of a collection express, Indicates the first Taiwan IoUT device, and All are integers greater than 1.
[0026] Without loss of generality, it is assumed that the target underwater region has fixed boundaries. , None of the Taiwanese AUVs were able to leave the target underwater area. , , and These represent the maximum length of the three dimensions respectively.
[0027] For simplicity, let's assume Taiwan AUV and Each IoUT device is located at the same fixed height, and the three-dimensional coordinates of the USV are represented as follows: , The three-dimensional coordinates are represented as , The three-dimensional coordinates are represented as , Indicates the altitude of the USV. express At a fixed height where the AUV is located, express The fixed height at which the IoUT device is located. For a given task execution time, it can be discretized into A time slot of equal length, Each time slot has a fixed duration. In addition, each A series of consecutive time slots are combined into a time period. , .
[0028] The working principle of a layered collaborative architecture is as follows. Initially, All AUVs are deployed at USV locations, and all are fully charged (i.e., ensuring remaining energy is...). ). The IoUT device operates at a fixed sampling frequency. Starting with underwater environment sensing data, initial data volume A value of 0 indicates that the USV possesses coarse-grained global information about the target underwater area, which is calculated every [time period]. Each time slot updates the global task allocation strategy in a centralized manner, based on... Weighted AoI of IoUT devices and The status of Taiwan AUV is as follows: A UV is assigned to the IoUT device to be collected. The AUV takes into account the high-level task allocation strategy and its own status to determine the final target IoUT device at the next moment, and drives from its current position to the target IoUT device to collect data.
[0029] In each time slot, the AUV sequentially performs movement, hovering data acquisition, data processing, and transmission operations. The data acquired from the target IoUT device is initially processed at the AUV and then transmitted to the USV via underwater acoustic communication for final calculation. The USV then transmits the calculation results to the shore base station.
[0030] Step 12: Construct the communication model, data acquisition model, and AoSI model. The communication model represents the underwater acoustic communication between each IoUT device and each AUV, the underwater acoustic communication between each AUV and the USV, and the radio frequency communication between the USV and the shore base station. The data acquisition model represents the data acquisition performed by each AUV from each IoUT device. The AoSI model is used to measure the timeliness and semantic value of data acquisition.
[0031] In practical implementation, for a layered collaborative architecture, we need to construct three models to describe it: a communication model, a data acquisition model, and an AoSI model. The communication model represents the underwater acoustic communication between each IoUT device and each AUV, the underwater acoustic communication between each AUV and the USV, and the radio frequency communication between the USV and the shore-based base station. The data acquisition model represents the data acquisition performed by each AUV from each IoUT device. The AoSI model is used to measure the timeliness and semantic value of data acquisition.
[0032] In one example, the communication model considers three phases of communication: underwater acoustic communication between each IoUT device and each AUV, underwater acoustic communication between each AUV and the USV, and radio frequency communication between the USV and the shore-based base station.
[0033] In one example, the data acquisition model represents a sequence of time periods. Within the time slot, each AUV collects data from its respective IoUT device. Due to limitations in communication bandwidth and operational scope, each AUV can only collect data from a maximum of one IoUT device at a time. , from Data is collected at the location. from The data transmission rate of the collected data is , Total amount of data collected for: ; in, express From Data is collected at the location. express Total data volume.
[0034] To comprehensively measure the timeliness and semantic value of underwater sensing data, we re-examine the timeliness of information from the perspective of "semantic value" and propose a new performance index, namely AoSI, which deeply couples time freshness with semantic information value.
[0035] The AoSI model is used to measure the timeliness and semantic value of data collection. In the time slot The AoSI is represented as , , for In the time slot AoI, for In the time slot Dynamic semantic weights are used when complex environmental changes with high entropy occur in the target underwater region. Increase, therefore The rapid ascent forces each AUV to prioritize data collection from the target underwater area. The core significance of AoSI lies not only in penalizing update delays, but also in imposing a more severe weighted penalty on delays in high-value information.
[0036] Based on the AoSI model, we transform the global optimization objective of multi-AUV collaborative data acquisition into minimizing the time-limited optimization within the task cycle. The global average AoSI is expressed as follows:
[0037] in, This represents the global average AoSI.
[0038] Step 13: Define the optimization problem based on the hierarchical collaboration architecture, communication model, data acquisition model and AoSI model, and model the optimization problem using Markov decision process.
[0039] In practical implementation, after constructing the hierarchical collaboration architecture, communication model, data acquisition model, and AoSI model, the optimization problem can be defined based on the hierarchical collaboration architecture, communication model, data acquisition model, and AoSI model, and the optimization problem can be modeled using Markov decision processes.
[0040] In our multi-AUV-assisted hierarchical data acquisition scenario, USVs and AUVs collaborate hierarchically across two different time scales to acquire sensing data of the underwater target's working area and transmit the processed data back to the shore-based base station. Within a given mission cycle... Within this framework, the data acquisition ratio is defined as follows: , ; in, express The total amount of sensory data remaining at the end of the task. This refers to the data collection ratio.
[0041] One of the optimization goals of optimization problems is to maximize the perceived value by optimizing the perception strategy. However, this could lead to an unfair data collection process, with some IoUT devices being collected too little or never. Therefore, it is necessary to consider the average perceived coverage of all IoUT devices. , Defined as: .
[0042] Next, in order to measure the efficiency of all AUVs performing sensing and computing tasks, it is necessary to consider improving the average energy efficiency of AUVs. , Defined as: ; in, This indicates the battery level of the AUV after it is fully charged.
[0043] Define the optimization problem, considering a heterogeneous multi-agent system (a single USV, a single agent ... A hierarchical data acquisition system consisting of AUVs (autonomous vehicles, air-to-ground vehicles, and air-to-ground vehicles) The IoUT device is positioned in a fixed underwater location. Within a limited timeframe... Inside, A single IoUT device continuously senses data from underwater at a fixed sampling frequency, USV and The AUVs from Taiwan collaborated to perform data acquisition tasks, aiming to maximize the efficiency of data collection within the constraints of the drones' limited energy. and and minimize The optimization problem is expressed as: ; ; Among them, constraints and constraints To ensure that each AUV can be assigned to only one IoUT device in each time slot, multiple AUVs can simultaneously collaborate on data acquisition for the same IoUT device, thus constraining... This indicates that the AUV can only operate within the target underwater area, ensuring that the AUV can efficiently collect data within the designated target underwater area, thus constraining... Ensure the distance between the two AUVs is greater than the preset safe distance. This prevents collisions and enables safe navigation, thus constraining... To ensure that the total power consumption of each AUV during the mission cycle is less than its maximum power consumption, and to ensure that the AUV operates efficiently without depleting its energy reserves, this constraint is necessary. Ensure AUV coverage of the entire target area. This represents the minimum average perceived coverage.
[0044] To effectively address the defined multi-AUV collaborative data acquisition optimization problem, we formalize it as a hierarchical Markov decision process (H-MDP), constructing MDP models for the two independent sub-problems of "high-level task allocation" and "low-level task execution." The high-level decision-maker (USV) performs global task allocation on a long-term scale, while the low-level decision-makers (each AUV) execute distributed motion control on a short-term scale.
[0045] For high-level MDPs, we model the high-level decision-making process as tuples. , For high-level global state space, To provide room for maneuver for higher-ups, For high-level reward functions, For higher-level state transition functions, This is a discount factor for higher-level employees.
[0046] High-level global state space In the time period High-level status It provides global information, including the status of the AUV cluster and the status of the IoUT devices. , Defined as: ; ; in, and They represent In time period Initial position information and remaining energy information, express Location, , and They represent In time period Initial AoI, remaining data volume, and dynamic semantic value weights.
[0047] High-level action space It is a joint task assignment vector, in each time period USV is Specify a target IoUT device. Defined as: ; in, express In time period Assigned to collect The data, and express In time period No tasks were assigned to them.
[0048] High-level reward function The high-level rewards aim to optimize long-term global goals, namely maximizing data collection ratio and geographical fairness, and minimizing average weighted AoI and total AUV energy consumption over a given time period. At the end, the reward for: ; , , ; in, , , , These are preset reward weighting coefficients used to balance the relative importance of various rewards. and These represent the changes in global data collection rate and geographical equity, respectively. This shows the change in the global average AoSI. Indicates the entire time period Inside Total energy consumed This represents the penalty for AUV energy depletion caused by infeasible allocation.
[0049] For the underlying Dec-POMDP, we model the underlying decision process as a distributed partially observable Markov decision process (Dec-POMDP), defined as a tuple. , It is a set of intelligent agents, i.e., an AUV set. For the underlying global state space, For joint action space, This is the underlying state transition function. For the underlying reward function, For joint observation space, For the observation function, This is the underlying discount factor.
[0050] Joint observation space In each time slot , It can only observe local states within its observable range, defined as follows: , Local observation Defined as: ; in, for One's own state of motion, express The remaining energy, for The set of neighboring AUVs within the observable range, Then it is Assigned Location information.
[0051] Joint Action Space When the IoUT device is within the communication range of the AUV, data acquisition will be performed automatically. The AUV only needs to control its movement. Therefore, in each time slot... , action Defined as: ; in, and They represent The changes in acceleration and angular velocity are discretized to meet the requirements of value decomposition-based multi-agent reinforcement learning for a discrete action space. and From the preset discrete sets respectively , By selecting corresponding values, a finite discrete action space is formed. and These represent the preset acceleration and angular velocity, respectively.
[0052] Underlying reward function The underlying rewards guide AUVs to complete their assigned tasks efficiently and collaboratively, ensuring safety and alignment with higher-level objectives. Defined as: ; , , ; ; in, This indicates a navigation reward to encourage... To it move, This indicates a data collection reward, which incentivizes the AUV for successfully collecting data. This represents an energy reward; the more energy consumed, the greater the penalty, thus incentivizing AUVs to learn more energy-efficient movement strategies. This indicates a safety bonus when an AUV collides with another AUV or goes over the edge. Punishment will be given at that time.
[0053] Step 14: The modeling results of the optimization problem are processed using the HRL-GIC hierarchical reinforcement learning algorithm to obtain the global optimal policy. The hierarchical collaborative architecture is then instructed to perform underwater data acquisition based on the optimal task allocation scheme and the optimal task execution scheme in the global optimal policy.
[0054] In practical implementation, after modeling the optimization problem, the HRL-GIC hierarchical reinforcement learning algorithm can be used to process the modeling results to obtain the global optimal policy. This policy instructs the hierarchical collaborative architecture to perform underwater data acquisition based on the optimal task allocation scheme and the optimal task execution scheme in the global optimal policy.
[0055] In one example, the principle of the HRL-GIC hierarchical reinforcement learning algorithm is as follows: Figure 3 As shown, it is specifically divided into two parts: high-level decision-making and low-level execution.
[0056] During the high-level decision-making cycle The goal of USV is to find the optimal task allocation scheme. In possession Taiwan AUV and In a scenario with a single IoUT device, the motion space size is... As the number of AUVs grows exponentially, the traditional DQN algorithm, which traverses all possible combinations of actions to compute... The previous approach is not feasible. To address this combinatorial optimization challenge, we introduce the concept of amortized optimization, using more efficient forward propagation in parametric models (such as neural networks) to replace the expensive process of iterative optimization.
[0057] Specifically, in addition to maintaining a global Q-network to evaluate the value of actions... In addition, an additional proposal network is trained. , Its purpose is to learn a probability distribution that can directly generate high-value candidate actions, thereby expanding the search space from the entire set. Reduced to a very small sample candidate set Through training To approximate the optimal action distribution, the computational cost is amortized into the network's weight updates. Among these... for The parameters, for The parameters.
[0058] In multi-AUV collaborative data acquisition tasks, there are structured dependencies between task allocations. For example, to avoid resource conflicts or duplicate coverage, if... I was assigned to So, Visiting the same location repeatedly should be avoided as much as possible. Therefore, in order to capture this conditional dependency, a proposal network will be used. The model is in the form of autoregressive decomposition: ; in, Indicates the preceding The allocation of AUVs already generated in Taiwan .
[0059] Subtasks are dynamically assigned to each AUV using an iterative approach. The allocation of AUVs depends not only on the global state. It also depends on the previous conditions The allocation of AUVs in Taiwan has been finalized.
[0060] Employing a GRU-based chained generative architecture, the network input includes not only global state embeddings. It also includes the embedding vector of the preceding action, in the first... When generating a step, GRU outputs information about the current hidden state based on the current hidden state. The Softmax probability distribution of the GRU algorithm, with its hidden states serving as the allocation history. The latent representation implicitly encodes the task space distribution and remaining resource requirements of the preceding AUV, and generates the result based on this condition. The proposed network ensures that the current allocation is spatially coordinated with previous decisions, thereby proactively avoiding potential task conflicts (such as duplicate allocation or resource waste) during the generation process.
[0061] To achieve more efficient generative scheduling, we learn the proposal network by first constructing a set of candidate assignment actions at each decision step. It contains from the proposed distribution Mid-sampling Each proposed action and samples from a uniform distribution of all globally assigned actions. For each exploration action, to evaluate these assigned actions, a global action value network is maintained, represented as... Based on the current action value function In the candidate set Search and determine the optimal assignment action with the highest Q value. , This process uses sampling approximation to replace the originally infeasible global maximization operation, effectively reducing the dimensionality of the combinatorial optimization problem.
[0062] Select the optimal allocation action Afterwards, we proposed the network. The goal of the proposed network is to learn how to assign, that is, to generate a distribution that approximates the currently known optimal assignment action as closely as possible through supervised learning. We update the parameters by minimizing the regularization loss function: ; The first term maximizes the log-likelihood of the optimal action assignment. Minimizing the first term makes the sample with the highest Q value more likely to appear under the proposal. The second term is the entropy regularization term of the proposal distribution, used to prevent the proposal distribution from converging prematurely to a deterministic policy, thereby maintaining long-term exploration capability. Through this update, the proposal distribution will gradually concentrate on the high Q-value action region, thus providing more candidate action samples in subsequent iterations.
[0063] Finally, we discussed the action value network. The update process, similar to that of a standard DQN network, aims to minimize the TD error. However, to accommodate the large action space, we need to compute the target value... At that time, the original global maximization factor will be... Replace with in the next time period sampling set Maximizing the above. We sample transfer data in the high-level experience replay pool. , Update loss function for: ; in, The target network is defined as [the network itself]. This design ensures that the accuracy of the value assessment gradually converges as the quality of the proposed network improves.
[0064] In the lower-level execution phase, AUVs need to efficiently execute tasks assigned at higher levels and handle multi-agent interactions under conditions of limited underwater acoustic communication bandwidth and partially observable environment.
[0065] Traditional MARL methods typically input all state information into a single network, making it difficult to maintain policy robustness under communication interruptions or high latency. To address this issue, we implement an action-value function for each AUV. Explicitly decomposed into independent Q values With collaborative Q value The product form. This design aims to dynamically balance the relationship between independent execution of high-level assignments and adaptive neighbor collaboration, in conjunction with an intent-driven communication mechanism to enhance the robustness of AUV mission execution in underwater environments with weak communication.
[0066] When communication is good It adjusts action choices to optimize group behavior. When communication is hindered, collaborative items degenerate, naturally and smoothly transitioning to independent execution mode to ensure uninterrupted task execution. The formal definition is as follows: ; in, Depends on the internal concealment state of the AUV , Includes current local observations Allocation with senior management , This represents the AUV's original intention to execute higher-level instructions without considering neighbor influences, and is used to measure the action. To complete the tasks assigned by higher management Independent execution contribution. Following communication, The estimated value used to weight each action is calculated by the Q-network. At the same time, he remained hidden. Information with neighbors (by the adaptive coordinator) The network takes filtered messages as input and evaluates the cooperative benefit of the current action in a multi-agent environment. The output layer uses a sigmoid activation function, with the cooperative Q-value acting as a modulating factor to dynamically adjust the base task value. The AUV ultimately selects the function that maximizes the action value. The action is executed, that is... .
[0067] To calculate the Q value and support the neighbor's collaborative strategy, we designed an intent-driven collaborative mechanism. The AUV needs to generate and broadcast its own action intent and establish information interaction with other neighboring AUVs. First, it is necessary to clarify the allocation of senior management based on one's own situation. To address the execution bias, we utilize GRU to maintain a fusion of local observations for AUVs. Allocation with senior management The hidden state, i.e. Local policy network computation And generate action intent , This represents the optimal action of the AUV when external cooperation is ignored. Subsequently, to adapt to the weak underwater communication environment, we encode the features of the hidden state. Encapsulate the action intent into a communication message , It is broadcast via an underwater acoustic channel. Compared to traditional methods that directly share raw observations, the transmitted encoded hidden state allows neighbors to implicitly perceive the local historical information and mission objectives.
[0068] AUV receives neighbor message Next, key interaction information needs to be extracted from the noisy communication environment. We introduce an adaptive coordinator module, which generates a cooperative mask for filtering incoming messages to determine the relevance of other AUVs to the current AUV's own intent. Definition For time slots Depend on The global message set sent by the AUV. By applying a co-mask for The obtained filtered message set, whose mask is calculated by the coordinator module, will be used to filter local messages. With all received neighbor messages Concatenate the data to construct an interactive feature sequence. , Then we will use the co-mask. The model is as follows: ; in, For coordinator network parameters, in a specific network architecture, It consists of a BiGRU followed by an MLP. The BiGRU is used to aggregate sequences. The bidirectional context information ensures that the weight evaluation of any neighbor depends not only on its own messages, but also on the global state of all other neighbors.
[0069] The generated mask is used to perform weighted filtering of neighbor messages to obtain a denoised message set. .
[0070] .
[0071] Based on this, the cooperative Q value is calculated. .
[0072] To ensure the effectiveness of the above mechanism, we designed a specific learning strategy for the underlying layer. To guarantee the consistency between the individual greedy policy and the globally optimal policy, we used the QMIX architecture to train the agent's Q-policy network. During the training phase, all AUVs were processed through a central hybrid network. The nonlinear combination of values is the global joint action value. The central hybrid network generates weights through a supernetwork and enforces monotonicity constraints. ,parameter Update by minimizing the global TD error: ; ; in, For the target value, and These are the parameters of the online policy network and the target policy network, respectively. The parameters of the target network are periodically changed from... copy.
[0073] This embodiment proposes an AoI-driven AUV-assisted high-efficiency data acquisition method for underwater IoT. Through innovative architecture design and algorithm optimization, it specifically addresses multiple core pain points in underwater IoT data acquisition, bringing significant benefits in multiple dimensions.
[0074] In terms of energy consumption optimization and device battery life, the solution adopts a spatiotemporally decoupled layered collaborative architecture that coordinates task allocation by USV and performs data collection by AUV, replacing the traditional multi-hop routing and fixed cluster head mode. This not only avoids the extra energy consumption caused by multi-hop transmission, but also reduces the redundant energy consumption caused by AUV preset trajectories through global optimal task allocation, effectively extending the operating life of IoUT devices powered by batteries without charging capabilities.
[0075] In terms of data timeliness and value assurance, the specially constructed AoSI model can accurately measure the timeliness and semantic value of data collection. Combined with the multi-AUV collaborative scheduling strategy, it can prioritize access (collection) of devices with large backlogs of stored data and high environmental data generation rates, thereby avoiding the problem of old data being overwritten due to buffer overflows. At the same time, for time-sensitive tasks such as underwater environmental monitoring and reconnaissance, it minimizes AoI and eliminates the collection of outdated and invalid information, ensuring the actual application value of the data.
[0076] In terms of adaptability to complex environments and scheduling efficiency, this embodiment integrates underwater acoustic communication and radio frequency communication models, which fully adapts to the communication characteristics of complex underwater environments. It makes up for the shortcomings of low bandwidth and long latency of single underwater acoustic communication. The optimization problem is modeled through Markov decision process, and the global optimal strategy is solved by HRL-GIC hierarchical reinforcement learning algorithm. It gets rid of the dependence of traditional multi-AUV scheduling on a large amount of prior environmental data, has strong online adaptability, and can efficiently cope with high-dimensional multi-objective optimization problems under multiple constraints such as turbulent environment, communication limitations, and dynamic changes in equipment status, and meet the dynamic upload requirements of large-scale IoUT devices.
[0077] In terms of network stability, replacing the traditional cluster head centralized collection mode with AUV mobile acquisition completely solves the problem of premature failure of the cluster head due to excessive communication overhead, significantly reduces the risk of potential network failures, and ensures the continuous reliability of the IoUT data acquisition process.
[0078] The steps described above are merely for clarity in describing the technical solution. In actual implementation, they can be combined into one step, or certain steps can be broken down into multiple steps, as long as they involve the same logical relationship, they are all within the scope of protection of this application. Any insignificant modifications or designs added to the algorithm or process, as long as they do not change the core of the algorithm or process, are also within the scope of protection of this application.
[0079] Another embodiment of this application proposes an AoI-driven AUV-assisted underwater IoT high-efficiency data acquisition system, implemented based on the AoI-driven AUV-assisted underwater IoT high-efficiency data acquisition method described in the above method embodiments. The details of the AoI-driven AUV-assisted underwater IoT high-efficiency data acquisition system proposed in this embodiment are described below. The following implementation details are provided for ease of understanding and are not essential for implementing this solution.
[0080] The specific structure of the AoI-driven AUV-assisted underwater IoT high-efficiency data acquisition system proposed in this embodiment is as follows: Figure 4 As shown, it includes: a hierarchical collaboration architecture building module 21, a model building module 22, an optimization problem modeling module 23, an optimization search module 24, and an execution module 25.
[0081] The hierarchical collaboration architecture construction module 21 is used to construct a spatiotemporally decoupled hierarchical collaboration architecture composed of multiple agents for underwater data acquisition scenarios. In the hierarchical collaboration architecture, the USV acts as the high-level decision-maker, responsible for high-level task allocation, while each AUV acts as the low-level decision-maker, responsible for low-level task execution.
[0082] Model building module 22 is used to build a communication model, a data acquisition model, and an AoSI model. The communication model represents the underwater acoustic communication between each IoUT device and each AUV, the underwater acoustic communication between each AUV and the USV, and the radio frequency communication between the USV and the shore base station. The data acquisition model represents the data acquisition performed by each AUV from each IoUT device. The AoSI model is used to measure the timeliness and semantic value of data acquisition.
[0083] The optimization problem modeling module 23 is used to define optimization problems based on a hierarchical collaborative architecture, communication model, data acquisition model and AoSI model, and to model the optimization problems using Markov decision processes.
[0084] The search module 24 is optimized to process the modeling results of the optimization problem using the HRL-GIC hierarchical reinforcement learning algorithm to obtain the globally optimal policy.
[0085] Module 25 is used to instruct the hierarchical collaborative architecture to perform underwater data acquisition based on the optimal task allocation scheme and the optimal task execution scheme in the global optimal strategy.
[0086] It is worth noting that all modules involved in this embodiment are logical modules. In practical applications, a logical module can be a physical module, a part of a physical module, or an organic combination of multiple physical modules. Furthermore, to highlight the innovative aspects of this application, this embodiment does not introduce modules that are not closely related to solving the technical problems proposed in this application. However, this does not mean that other modules are absent from this embodiment.
[0087] It is not difficult to see that this embodiment is a system embodiment corresponding to the above method embodiments, and this embodiment can be implemented in conjunction with the above method embodiments. The relevant technical details and technical effects mentioned in the above method embodiments are still valid in this embodiment, and will not be repeated here to reduce repetition. Accordingly, the relevant technical details mentioned in this embodiment can also be applied to the above method embodiments.
[0088] Another embodiment of this application provides an electronic device, such as Figure 5 As shown, it includes a processor 31 and a memory 32. The memory 32 stores instructions that the processor 31 can execute. When the processor 31 is configured to execute the instructions, the electronic device can realize an AoI-driven AUV-assisted underwater IoT high-efficiency data acquisition method as described in the above method embodiment.
[0089] The memory and processor are connected via a bus, which includes any number of interconnecting buses and bridges, connecting various circuits of one or more processors and the memory. The bus can also connect various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and will not be described further herein. The bus interface provides an interface between the bus and the transceiver. The transceiver can be a single component or multiple components, such as multiple receivers and transmitters, providing a unit for communicating with various other devices over a transmission medium. Data processed by the processor is transmitted over the wireless medium via an antenna, which further receives data and transmits it to the processor.
[0090] The processor manages the bus and general processing, and also provides various functions, including timing, peripheral interfaces, voltage regulation, power management, and other control functions. Memory is used to store data used by the processor during operation.
[0091] Another embodiment of this application proposes a computer-readable storage medium storing a computer program that, when executed by a processor, can implement an AoI-driven AUV-assisted underwater IoT high-efficiency data acquisition method as described in the above method embodiments.
[0092] That is, those skilled in the art will understand that all or part of the steps in the above method embodiments can be implemented by a program instructing related hardware. The program is stored in a storage medium and includes several instructions to cause a device (such as a microcontroller, chip, etc.) or processor to execute all or part of the steps of the method described in the method embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory, random access memory, magnetic disks, or optical disks.
[0093] It will be understood by those skilled in the art that the above embodiments are specific implementations of this application, and various changes in form and detail can be made in practical applications without departing from the spirit and scope of this application. For those skilled in the art, several improvements and modifications can be made without departing from the principles of this application, and these improvements and modifications are also considered to be within the scope of protection of this application.
Claims
1. An AoI-driven AUV-assisted underwater IoT high-efficiency data acquisition method, implemented based on a spatiotemporal decoupling layered collaborative architecture, characterized in that, The method includes: For underwater data acquisition scenarios, a spatiotemporally decoupled hierarchical collaborative architecture composed of multiple agents is constructed. In this hierarchical collaborative architecture, the USV acts as the high-level decision-maker, responsible for high-level task allocation, while each AUV acts as the low-level decision-maker, responsible for low-level task execution. A communication model, a data acquisition model, and an AoSI model are constructed. The communication model represents the underwater acoustic communication between each IoUT device and each AUV, the underwater acoustic communication between each AUV and the USV, and the radio frequency communication between the USV and the shore base station. The data acquisition model represents the data acquisition performed by each AUV from each IoUT device. The AoSI model is used to measure the timeliness and semantic value of data acquisition. The optimization problem is defined based on a hierarchical collaborative architecture, communication model, data acquisition model, and AoSI model, and the optimization problem is modeled using a Markov decision process. The HRL-GIC hierarchical reinforcement learning algorithm is used to process the modeling results of the optimization problem to obtain the global optimal policy. This policy instructs the hierarchical collaborative architecture to perform underwater data acquisition based on the optimal task allocation scheme and the optimal task execution scheme in the global optimal policy.
2. The AoI-driven AUV-assisted high-efficiency data acquisition method for underwater IoT as described in claim 1, characterized in that, In underwater data acquisition scenarios, multiple AUVs continuously cruise in turbulent ocean environments and determine the target IoUT devices that need to be accessed for data acquisition based on the status information of the IoUT devices broadcast by the USVs. Assume there is a shore-based base station and a USV. Taiwan AUV and Taiwan IoUT device, Taiwan AUV is a collection express, Indicates the first Taiwanese AUV, A single IoUT device consists of a collection express, Indicates the first Taiwan IoUT device, and All integers are greater than 1, assuming the target underwater region has fixed boundaries. , None of the Taiwanese AUVs were able to leave the target underwater area. , , and These represent the maximum length of the three dimensions respectively; Assumption Taiwan AUV and Each IoUT device is located at the same fixed height, and the three-dimensional coordinates of the USV are represented as follows: , The three-dimensional coordinates are represented as , The three-dimensional coordinates are represented as , Indicates the altitude of the USV. express At a fixed height where the AUV is located, express The fixed height at which the IoUT device is located. For a given task execution time, it can be discretized into A time slot of equal length, Each time slot has a fixed duration. ,Every A series of consecutive time slots are combined into a time period. , . At the initial moment, All AUVs were deployed at USV locations, and all were fully charged. The IoUT device operates at a fixed sampling frequency. Starting with underwater environment sensing data, initial data volume A value of 0 indicates that the USV possesses coarse-grained global information about the target underwater area, which is calculated every [time period]. Each time slot updates the global task allocation strategy in a centralized manner, based on... Weighted AoI of IoUT devices and The status of Taiwan AUV is as follows: A UV is assigned to the IoUT device to be collected. The AUV comprehensively considers the high-level task allocation strategy and its own status to determine the final target IoUT device at the next moment, and drives from its current position to the target IoUT device to collect data; In each time slot, the AUV sequentially performs movement, hovering data acquisition, data processing, and transmission operations. The data acquired from the target IoUT device is initially processed at the AUV and then transmitted to the USV via underwater acoustic communication for final calculation. The USV then transmits the calculation results to the shore base station.
3. The AoI-driven AUV-assisted underwater IoT high-efficiency data acquisition method according to claim 2, characterized in that, The communication model considers three stages of communication: underwater acoustic communication between each IoUT device and each AUV, underwater acoustic communication between each AUV and USV, and radio frequency communication between USV and shore base station. The data acquisition model represents a data collection process within an ordered time period. Within the time slot, each AUV collects data from its respective IoUT device. Due to limitations in communication bandwidth and operational scope, each AUV can only collect data from a maximum of one IoUT device at a time. , from Data is collected at the location. from The data transmission rate of the collected data is , Total amount of data collected for: ; in, express From Data is collected at the location. express Total data volume; The AoSI model is used to measure the timeliness and semantic value of data collection. In the time slot The AoSI is represented as , , for In the time slot AoI, for In the time slot Dynamic semantic weights are used when complex environmental changes with high entropy occur in the target underwater region. Increase, therefore The rapid ascent forces each AUV to prioritize collecting data from the target underwater area; Based on the AoSI model, the global optimization objective of multi-AUV collaborative data acquisition is transformed into minimizing the time required for the task cycle. The global average AoSI is expressed as follows: in, This represents the global average AoSI.
4. The AoI-driven AUV-assisted underwater IoT high-efficiency data acquisition method according to claim 3, characterized in that, The optimization problem is defined based on the hierarchical collaboration architecture, communication model, data acquisition model, and AoSI model, including: In a given task cycle Within this framework, the data acquisition ratio is defined as follows: , ; in, express The total amount of sensory data remaining at the end of the task. For data collection ratio; One of the optimization goals of optimization problems is to maximize the perceived value by optimizing the perception strategy. However, this could lead to an unfair data collection process, with some IoUT devices being collected too little or never. Therefore, it is necessary to consider the average perceived coverage of all IoUT devices. , Defined as: ; To measure the efficiency of all AUVs performing sensing and computing tasks, it is necessary to consider improving the average energy efficiency of AUVs. , Defined as: ; in, This indicates the battery level of the AUV after it is fully charged; Define the optimization problem within a finite time frame. Inside, A single IoUT device continuously senses data from underwater at a fixed sampling frequency, USV and The AUVs from Taiwan collaborated to perform data acquisition tasks, aiming to maximize the efficiency of data collection within the constraints of the drones' limited energy. and and minimize The optimization problem is expressed as: ; ; Among them, constraints and constraints To ensure that each AUV can be assigned to only one IoUT device in each time slot, multiple AUVs can simultaneously collaborate on data acquisition for the same IoUT device, thus constraining... This indicates that the AUV can only operate within the target underwater area, ensuring that the AUV can efficiently collect data within the designated target underwater area, thus constraining... Ensure the distance between the two AUVs is greater than the preset safe distance. This prevents collisions and enables safe navigation, thus constraining... To ensure that the total power consumption of each AUV during the mission cycle is less than its maximum power consumption, and to ensure that the AUV operates efficiently without depleting its energy reserves, this constraint is necessary. Ensure AUV coverage of the entire target area. This represents the minimum average perceived coverage.
5. The AoI-driven AUV-assisted underwater IoT high-efficiency data acquisition method according to claim 4, characterized in that, Modeling optimization problems using Markov decision processes includes: MDP models are constructed for high-level task allocation and low-level task execution, respectively. High-level decision-makers allocate global tasks on a long-term scale, while low-level decision-makers execute distributed motion control on a short-term scale. Modeling high-level decision-making processes as tuples , For high-level global state space, To provide room for maneuver for higher-ups, For high-level reward functions, For higher-level state transition functions, For high-level discount factors; High-level global state space In the time period High-level status It provides global information, including the status of the AUV cluster and the status of the IoUT devices. , Defined as: ; ; in, and They represent In time period Initial position information and remaining energy information, express Location, , and They represent In time period Initial AoI, remaining data volume, and dynamic semantic value weights; High-level action space It is a joint task assignment vector, in each time period USV is Specify a target IoUT device. Defined as: ; in, express In time period Assigned to collect The data, and express In time period No task was assigned to him / her. High-level reward function The high-level rewards aim to optimize long-term global goals, namely maximizing data collection ratio and geographical fairness, and minimizing average weighted AoI and total AUV energy consumption over a given time period. At the end, the reward for: ; , , ; in, , , , These are preset reward weighting coefficients used to balance the relative importance of various rewards. and These represent the changes in global data collection rate and geographical equity, respectively. This shows the change in the global average AoSI. Indicates the entire time period Inside Total energy consumed This represents the penalty for AUV energy depletion caused by infeasible allocation; The underlying decision-making process is modeled as a distributed, partially observable Markov decision process, defined as a tuple. , For the underlying global state space, For joint action space, This is the underlying state transition function. For the underlying reward function, For joint observation space, For the observation function, The underlying discount factor; Joint observation space In each time slot , It can only observe local states within its observable range, defined as follows: , Local observation Defined as: ; in, for One's own state of motion, express The remaining energy, for The set of neighboring AUVs within the observable range, Then it is Assigned Location information; Joint Action Space When the IoUT device is within the communication range of the AUV, data acquisition will be performed automatically. The AUV only needs to control its movement. Therefore, in each time slot... , action Defined as: ; in, and They represent The changes in acceleration and angular velocity are discretized to meet the requirements of value decomposition-based multi-agent reinforcement learning for a discrete action space. and From the preset discrete sets respectively , By selecting corresponding values, a finite discrete action space is formed. and These represent the preset acceleration and angular velocity, respectively. Underlying reward function The underlying rewards guide AUVs to complete their assigned tasks efficiently and collaboratively, ensuring safety and alignment with higher-level objectives. Defined as: ; , , ; ; in, This indicates a navigation reward to encourage... To it move, This indicates a data collection reward, which incentivizes the AUV for successfully collecting data. This represents an energy reward; the more energy consumed, the greater the penalty, thus incentivizing AUVs to learn more energy-efficient movement strategies. This indicates a safety bonus when an AUV collides with another AUV or goes over the edge. Punishment will be given at that time.
6. The AoI-driven AUV-assisted underwater IoT high-efficiency data acquisition method according to claim 5, characterized in that, The HRL-GIC hierarchical reinforcement learning algorithm is used to process the modeling results of the optimization problem to obtain the global optimal policy, which is divided into two parts: high-level decision-making and low-level execution. During the high-level decision-making cycle The goal of USV is to find the optimal task allocation scheme. In possession Taiwan AUV and In a scenario with a single IoUT device, the motion space size is... It introduces the idea of amortization optimization, using more efficient forward propagation in the parameterized model to replace iterative optimization; In addition to maintaining a global Q-network for evaluating the value of actions In addition, an additional proposal network is trained. , Its purpose is to learn a probability distribution that can directly generate high-value candidate actions, thereby expanding the search space from the entire set. Reduced to a very small sample candidate set Through training To approximate the optimal action distribution, the computational cost is amortized into the network's weight updates; among which, for The parameters, for Parameters; In multi-AUV collaborative data acquisition tasks, there are structured dependencies between task assignments. To capture these conditional dependencies, a proposal network is used. The model is in the form of autoregressive decomposition: ; in, Indicates the preceding The allocation of AUVs already generated in Taiwan ; Subtasks are dynamically assigned to each AUV using an iterative approach. The allocation of AUVs depends not only on the global state. It also depends on the previous conditions The allocation of AUVs in Taiwan has been finalized; Employing a GRU-based chained generative architecture, the network input includes not only global state embeddings. It also includes the embedding vector of the preceding action, in the first... When generating a step, GRU outputs information about the current hidden state based on the current hidden state. The Softmax probability distribution of the GRU algorithm, with its hidden states serving as the allocation history. The latent representation implicitly encodes the task space distribution and remaining resource requirements of the preceding AUV, and generates the result based on this condition. The proposed network ensures that the current allocation is spatially consistent with previous decisions, thereby proactively avoiding potential task conflicts during the generation process. The proposal network is learned by first constructing a set of candidate assignment actions at each decision step. It contains from the proposed distribution Mid-sampling Each proposed action and samples from a uniform distribution of all globally assigned actions. For each exploration action, to evaluate these assigned actions, a global action value network is maintained, represented as... Based on the current action value function In the candidate set Search and determine the optimal assignment action with the highest Q value. , ; Select the optimal allocation action Subsequently, regarding the proposed network The goal of the proposed network is to learn how to assign, that is, to generate a distribution that approximates the currently known optimal assignment action as closely as possible through supervised learning. The parameters are updated by minimizing the regularization loss function: ; The first term maximizes the log-likelihood of the optimal allocation action, while minimizing the first term makes the sample with the highest Q value more likely to appear under the proposal. The second term is the entropy regularization term of the proposal distribution, which is used to prevent the proposal distribution from converging to the deterministic policy too early, thereby maintaining long-term exploration capability. Finally, regarding the action value network The update process aims to minimize the TD error and, in order to accommodate the large action space, optimize the calculation of the target value. At that time, the original global maximization factor will be... Replace with in the next time period sampling set Maximizing the sampling and transfer of data in the high-level experience replay pool. , Update loss function for: ; in, For the target network.
7. The AoI-driven AUV-assisted underwater IoT high-efficiency data acquisition method according to claim 6, characterized in that, In the low-level execution phase, AUVs need to efficiently execute tasks assigned at higher levels and handle multi-agent interactions under conditions of limited underwater acoustic communication bandwidth and partially observable environment. The action value function of each AUV Explicitly decomposed into independent Q values With collaborative Q value The product form aims to dynamically balance the relationship between independent execution of high-level allocation and adaptive neighbor cooperation, combined with an intent-driven communication mechanism to improve the robustness of AUV mission execution in underwater environments with weak communication. When communication is good, It modulates action selection to optimize group behavior. When communication is blocked, cooperative items degenerate, naturally and smoothly transitioning to independent execution mode to ensure uninterrupted task execution. The formal definition is as follows: ; in, Depends on the internal concealment state of the AUV , Includes current local observations Allocation with senior management , This represents the AUV's original intention to execute higher-level instructions without considering neighbor influences, and is used to measure the action. To complete the tasks assigned by higher management Independent execution contribution, after communication, The estimated value used to weight each action is calculated by the Q-network. At the same time, he remained hidden. Information with neighbors The input layer evaluates the cooperative benefits of the current action in a multi-agent environment. The output layer uses a sigmoid activation function, with the cooperative Q-value as a modulating factor to dynamically adjust the value of the basic task. The AUV ultimately selects the function that maximizes the action value. The action is executed; AUVs need to generate and broadcast their own action intentions and establish information exchange with other neighboring AUVs. First, it is necessary to clarify the allocation of senior management based on one's own situation. The execution tendency is to use GRU to maintain a fusion of local observations for AUV. Allocation with senior management The hidden state, i.e. Local policy network computation And generate action intent , This represents the optimal action of the AUV when external cooperation is ignored. To adapt to the weak communication environment underwater, the hidden state features are encoded. Encapsulate the action intent into a communication message , And broadcast it via an underwater acoustic channel; AUV receives neighbor message Next, key interaction information needs to be extracted from the noisy communication environment. An adaptive coordinator module is introduced to generate a cooperative mask for filtering incoming messages, thereby determining the relevance of other AUVs to the current AUV's own intent. For time slots Depend on The global message set sent by the AUV. By applying a co-mask for The obtained filtered message set, whose mask is calculated by the coordinator module, will be used to filter local messages. With all received neighbor messages Concatenate the data to construct an interactive feature sequence. , Then use the co-mask The model is as follows: ; in, For coordinator network parameters, in a specific network architecture, It consists of a BiGRU followed by an MLP. The BiGRU is used to aggregate sequences. The bidirectional context information ensures that the weight evaluation of any neighbor depends not only on its own messages, but also on the global state of all other neighbors. The generated mask is used to perform weighted filtering of neighbor messages to obtain a denoised message set. : ; Based on this, the cooperative Q value is calculated. ; To ensure the effectiveness of the mechanism, a specific learning strategy is designed for the underlying layer. To ensure the consistency between the individual greedy strategy and the global optimal strategy, the Q-policy network of AUV is trained using the QMIX architecture. During the training phase, all AUV individuals are processed through a central hybrid network. The nonlinear combination of values is the global joint action value. The central hybrid network generates weights through a supernetwork and enforces monotonicity constraints. ,parameter Update by minimizing the global TD error: ; ; in, For the target value, and These are the parameters of the online policy network and the target policy network, respectively. The parameters of the target network are periodically changed from... copy.
8. An AoI-driven AUV-assisted underwater IoT high-efficiency data acquisition system, implemented based on the AoI-driven AUV-assisted underwater IoT high-efficiency data acquisition method as described in any one of claims 1 to 7, characterized in that, The system includes: The hierarchical collaboration architecture construction module is used to build a spatiotemporally decoupled hierarchical collaboration architecture composed of multiple agents for underwater data acquisition scenarios. In the hierarchical collaboration architecture, the USV acts as the high-level decision-maker, responsible for high-level task allocation, while each AUV acts as the low-level decision-maker, responsible for low-level task execution. The model building module is used to build communication models, data acquisition models, and AoSI models. The communication model represents the underwater acoustic communication between each IoUT device and each AUV, the underwater acoustic communication between each AUV and the USV, and the radio frequency communication between the USV and the shore base station. The data acquisition model represents the data acquisition performed by each AUV from each IoUT device. The AoSI model is used to measure the timeliness and semantic value of data acquisition. The optimization problem modeling module is used to define optimization problems based on a hierarchical collaborative architecture, communication model, data acquisition model, and AoSI model, and to model the optimization problems using Markov decision processes. An optimized search module is used to process the modeling results of the optimization problem using the HRL-GIC hierarchical reinforcement learning algorithm to obtain the globally optimal policy; The execution module is used to instruct the hierarchical collaborative architecture to perform underwater data acquisition based on the optimal task allocation scheme and the optimal task execution scheme in the global optimal strategy.
9. An electronic device, characterized in that, include: A processor and a memory, wherein the memory stores instructions executable by the processor, and the processor is configured to, when executing the instructions, enable the electronic device to implement an AoI-driven AUV-assisted underwater Internet of Things high-efficiency data acquisition method as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it can implement an AoI-driven AUV-assisted underwater Internet of Things high-efficiency data acquisition method as described in any one of claims 1 to 7.